Long document retrieval aims to fetch query-relevant documents from a large-scale collection, where knowledge distillation has become de facto to improve a retriever by mimicking a heterogeneous yet powerful cross-encoder. However, in contrast to passages or sentences, retrieval on long documents suffers from the scope hypothesis that a long document may cover multiple topics. This maximizes their structure heterogeneity and poses a granular-mismatch issue, leading to an inferior distillation efficacy. In this work, we propose a new learning framework, fine-grained distillation (FGD), for long-document retrievers. While preserving the conventional dense retrieval paradigm, it first produces global-consistent representations crossing different fine granularity and then applies multi-granular aligned distillation merely during training. In experiments, we evaluate our framework on two long-document retrieval benchmarks, which show state-of-the-art performance.
translated by 谷歌翻译
联合学习(FL)是一种机器学习范式,允许分散的客户在不共享其私人数据的情况下进行协作学习。但是,过度的计算和沟通要求对当前的FL框架构成挑战,尤其是在训练大型模型时。为了防止这些问题阻碍FL系统的部署,我们提出了一个轻巧的框架,客户共同学习融合由多个固定预训练的模型生成的表示形式,而不是从SCRATCH培训大型模型。这通过考虑如何从预先训练的模型中捕获更多特定于客户的信息,并共同提高每个客户利用这些现成模型的能力,从而导致我们解决了一个更实用的FL问题。在这项工作中,我们设计了一种联合原型对比度学习(FEDPCL)方法,该方法通过其类原型共享客户的知识,并以原型对比度方式构建特定于客户的表示。共享原型而不是可学习的模型参数可以使每个客户以个性化的方式融合表示表示,同时以紧凑的形式保持共享知识以进行有效的通信。我们在轻量级框架中对拟议的FEDPCL进行了彻底的评估,以测量和可视化其在流行的FL数据集上融合各种预训练模型的能力。
translated by 谷歌翻译
众所周知,由出色的文档级神经机器翻译(NMT)模型产生的翻译是一致且连贯的。但是,像BLEU这样的现有句子级评估指标几乎无法反映模型在文档级别的性能。为了解决这个问题,我们在本文中提出了一种话语凝聚评估方法(DCOEM),并贡献了一个新的测试套件,该套件考虑了四个凝聚力的方式(参考,连接,替代和词汇凝聚力),以衡量文档翻译的凝聚力。最近的文档级NMT系统的评估结果表明,我们的方法在估计文档级别的翻译方面是实用且至关重要的。
translated by 谷歌翻译
生物医学实体的因果关系提取是生物医学文本挖掘中最复杂的任务之一,涉及两种信息:实体关系和实体功能。一种可行的方法是将关系提取和功能检测作为两个独立的子任务。但是,这种单独的学习方法忽略了它们之间的内在相关性,并导致性能不令人满意。在本文中,我们提出了一个联合学习模型,该模型结合了实体关系提取和实体功能检测以利用其共同点并捕获其相互关系,以提高生物医学因果关系提取的性能。同时,在模型训练阶段,损失函数中的不同功能类型分配了不同的权重。具体而言,负功能实例的惩罚系数增加以有效提高功能检测的精度。 Biocreative-V轨道4语料库的实验结果表明,我们的联合学习模型在BEL语句提取中的表现优于单独的模型,在第2阶段和第1阶段评估中的测试集中,F1得分分别达到58.4%和37.3%。这表明,与其他系统相比,我们的联合学习系统达到了第2阶段的最新性能。
translated by 谷歌翻译
排名者在事实上的“检索和rerank”管道中起着必不可少的作用,但其训练仍然落后 - 从中​​度的负面因素或/和/和/和作为回收者的辅助模块中学习。在这项工作中,我们首先确定了强大的排名者的两个主要障碍,即是由训练有素的回猎犬和非理想的负面负面的固有标签噪声,该噪声是为高能力的排名所采样的。因此,我们提出多个检索器,因为负面发电机改善了排名者的鲁棒性,其中i)涉及广泛的分发标签噪声,使排名者与每个噪声分布相对,而ii)与排名相对较接近排名负分配,导致更具挑战性的培训。为了评估我们的强大排名者(称为r $^2 $ anker),我们在各种环境中进行了有关流行通道检索基准测试的各种实验,包括BM25级,全等级,回收者蒸馏等。经验结果验证了新的州 - 新州 - 新州 - 我们模型的效果。
translated by 谷歌翻译
知识共享和模型个性化是应对联邦学习(FL)中非IID挑战的重要组成部分。大多数现有的FL方法侧重于两个极端:1)学习共享模型,以使用非IID数据为所有客户提供服务,以及2)为每个客户(即个性化fl)学习个性化模型。有一个权衡解决方案,即群集或集群个性化的FL,旨在将相似的客户聚集到一个集群中,然后在集群中为所有客户学习共享模型。本文是通过将群集群集制定为可以统一现有方法的双层优化框架来重新审视群集的研究。我们提出了一个新的理论分析框架,以通过考虑客户之间的凝聚力来证明融合。此外,我们以一种称为加权聚类联合学习(WECFL)的算法体现了该框架。经验分析验证了理论结果,并证明了在拟议的集群非IID设置下提出的WECFL的有效性。
translated by 谷歌翻译
在联合学习(FL)中的客户端的异质性通常会在梯度空间中发生客户的知识聚合时阻碍优化融合和泛化性能。例如,客户端可以在数据分发,网络延迟,输入/输出空间和/或模型架构方面不同,这可以很容易地导致其本地梯度的未对准。为了提高异质性的容忍度,我们提出了一种新的联合原型学习(FedProto)框架,其中客户端和服务器传达了抽象类原型而不是梯度。 FEDPROTO聚合从不同客户端收集的本地原型,然后将全局原型发送回所有客户端,以规范本地模型的培训。每个客户端的训练旨在最大限度地减少本地数据上的分类错误,同时保持所产生的本地原型靠近相应的全球范围。此外,我们在非凸起目标下对FedProto的收敛速度提供了理论分析。在实验中,我们提出了一种针对异构FL定制的基准设置,FEDPROTO优于多个数据集上的几种方法。
translated by 谷歌翻译
接受场(RF)的大小一直是时间序列分类任务中一维卷积神经网络(1D-CNN)的最重要因素之一。已经采取了巨大的努力来选择适当的大小,因为它对性能产生了巨大影响,并且每个数据集都有很大的不同。在本文中,我们为1D-CNN提出了一个Omni级块(OS-Block),其中内核大小由简单而通用的规则决定。特别是,它是一组内核大小,可以根据时间序列的长度通过多个素数组成,可以有效地覆盖不同数据集的最佳RF大小。实验结果表明,具有OSBlock的模型可以达到与搜索最佳RF尺寸的模型相似的性能,并且由于最佳的最佳RF尺寸捕获能力,具有OS-Block的简单1D-CNN模型可实现最新状态。四个时间序列基准的ART性能,包括来自多个域的单变量和多元数据。全面的分析和讨论阐明了为什么OS-Block可以在不同数据集中捕获最佳的RF尺寸。可用代码[https://github.com/wensi-tang/os-cnn]
translated by 谷歌翻译
An enhanced geothermal system is essential to provide sustainable and long-term geothermal energy supplies and reduce carbon emissions. Optimal well-control scheme for effective heat extraction and improved heat sweep efficiency plays a significant role in geothermal development. However, the optimization performance of most existing optimization algorithms deteriorates as dimension increases. To solve this issue, a novel surrogate-assisted level-based learning evolutionary search algorithm (SLLES) is proposed for heat extraction optimization of enhanced geothermal system. SLLES consists of classifier-assisted level-based learning pre-screen part and local evolutionary search part. The cooperation of the two parts has realized the balance between the exploration and exploitation during the optimization process. After iteratively sampling from the design space, the robustness and effectiveness of the algorithm are proven to be improved significantly. To the best of our knowledge, the proposed algorithm holds state-of-the-art simulation-involved optimization framework. Comparative experiments have been conducted on benchmark functions, a two-dimensional fractured reservoir and a three-dimensional enhanced geothermal system. The proposed algorithm outperforms other five state-of-the-art surrogate-assisted algorithms on all selected benchmark functions. The results on the two heat extraction cases also demonstrate that SLLES can achieve superior optimization performance compared with traditional evolutionary algorithm and other surrogate-assisted algorithms. This work lays a solid basis for efficient geothermal extraction of enhanced geothermal system and sheds light on the model management strategies of data-driven optimization in the areas of energy exploitation.
translated by 谷歌翻译
Facial Expression Recognition (FER) in the wild is an extremely challenging task. Recently, some Vision Transformers (ViT) have been explored for FER, but most of them perform inferiorly compared to Convolutional Neural Networks (CNN). This is mainly because the new proposed modules are difficult to converge well from scratch due to lacking inductive bias and easy to focus on the occlusion and noisy areas. TransFER, a representative transformer-based method for FER, alleviates this with multi-branch attention dropping but brings excessive computations. On the contrary, we present two attentive pooling (AP) modules to pool noisy features directly. The AP modules include Attentive Patch Pooling (APP) and Attentive Token Pooling (ATP). They aim to guide the model to emphasize the most discriminative features while reducing the impacts of less relevant features. The proposed APP is employed to select the most informative patches on CNN features, and ATP discards unimportant tokens in ViT. Being simple to implement and without learnable parameters, the APP and ATP intuitively reduce the computational cost while boosting the performance by ONLY pursuing the most discriminative features. Qualitative results demonstrate the motivations and effectiveness of our attentive poolings. Besides, quantitative results on six in-the-wild datasets outperform other state-of-the-art methods.
translated by 谷歌翻译